Back

Statistics in Medicine

Wiley

Preprints posted in the last 7 days, ranked by how well they match Statistics in Medicine's content profile, based on 40 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit.

1
Bayesian Borrowing of External Information in Clinical Trials: A Comparison of MAP, RMAP, and SAM Priors

Choi, L.; McNeer, E.; Beck, C. A.; Neul, J. L.

2026-08-31 pharmacology and therapeutics 10.64898/2026.08.26.26360843 medRxiv
Top 0.1%
10.0%
Show abstract

Bayesian borrowing of external information can improve trial efficiency, particularly in pediatric and rare disease settings where patient populations are limited, but may introduce bias and inflate the Type~I error rate when the trial differs from external studies. Recent U.S. Food and Drug Administration (FDA) draft Bayesian guidance emphasizes careful evaluation of external information, prior specification, and assessment of operating characteristics. This paper compares three meta-analytic-predictive (MAP)-based methods for Bayesian borrowing: the MAP prior, robust MAP (RMAP) prior, and self-adapting mixture (SAM) prior. An adaptive platform trial design in Rett syndrome is used as a case study. Simulation studies evaluate frequentist operating characteristics under varying prior--data conflict, between-study heterogeneity, treatment effects, and clinically significant differences (CSDs) for the SAM prior. The MAP prior achieved the greatest efficiency when external and current data were compatible but exhibited the largest bias under substantial prior--data conflict. The RMAP priors improved robustness through fixed robust-component weights, whereas the SAM prior adaptively adjusted borrowing and was less sensitive to prior--data conflict while retaining efficiency gains when the data were compatible. Although the CSD influenced the degree of adaptive borrowing, as reflected by effective sample size, it had only a modest impact on frequentist operating characteristics. Sensitivity analyses using a skeptical robust component yielded similar qualitative conclusions, while accentuating the differences between the MAP and RMAP priors. These findings provide guidance for evaluating and selecting MAP-based borrowing strategies before trial implementation, particularly in rare disease settings, consistent with current FDA recommendations.

2
A Measurement-Based Care Strategy for Buprenorphine-Naloxone Treatment (Bup-MBC): Development of an EHR-Integrated Intervention

Reese, T.; Audet, C.; Ancker, J.; Wright, A.; Marcovitz, D.; Kast, K. A.; Bridges, J.; Tindle, H.; Shah, M.; von Horn, A.; Matheny, M. E.

2026-09-01 addiction medicine 10.64898/2026.08.27.26361539 medRxiv
Top 0.1%
4.3%
Show abstract

Introduction: Risk of recurrent opioid use during buprenorphine-naloxone (bup-nx) treatment is dynamic and remains elevated after initiation, with vulnerability shaped in part by treatment intensity and gaps between visits, yet routine outpatient care relies on episodic encounters and retrospective data. This mismatch can delay recognition of emerging instability and limit timely treatment adjustments. This paper reports the development and specification of an intervention strategy to address this mismatch. Methods: We used a structured, multi-phase design process to specify and configure a measurement-based care (MBC) strategy for bup-nx treatment (Bup-MBC) in outpatient addiction clinics through three phases: (1) a systematic review of patient-reported outcome measures (PROMs) for substance use treatment; (2) a qualitative needs assessment using the Theoretical Domains Framework and COM-B (Capability, Opportunity, Motivation-Behavior) model to identify gaps in risk monitoring, agency, and trust; and (3) iterative co-design with multidisciplinary clinicians to refine workflow fit and trust-preserving use of data. Patients informed item and feedback content during the needs assessment but did not participate in the co-design cycles. Results: Bup-MBC integrates (1) brief between-visit PROMs (e.g., withdrawal, craving, adherence); (2) immediate non-punitive patient feedback; (3) clinician-facing summaries and non-directive prompts in the electronic health record (EHR); and (4) an opt-in between-visit outreach pathway with predefined safety triggers, all configured within existing EHR and patient portal infrastructure. It targets patient and clinician capability to recognize changes in risk, opportunity for action through structured monitoring and visit preparation, and trust and agency through non-punitive communication, without adding substantial burden. The full measure set, severity bands, and question-to-action map are provided as supplementary material. Key trade-offs included prioritizing single-item measures for feasibility, balancing opt-in outreach with safety overrides, and assuming routine clinician use of summaries. Conclusion: This development study specifies an EHR-integrated MBC strategy for outpatient bup-nx treatment. As single-center design work with co-design limited to clinicians and delivery contingent on portal or text-message access, its outputs are hypotheses about mechanism and fit rather than demonstrated effects. Feasibility studies are needed to evaluate uptake, acceptability, workflow fit, and effects on treatment.

3
ICONIC: An R Package for Integrating Instrumental Variable- and Negative-Control-Informed Causal Discovery and Diagnostics in Multiomic Studies

Bresnahan, S. T.; Xiong, C.; Head, T.; Chang, Y.-H.; Bhattacharya, A.; Huang, J. Y.

2026-08-31 genetic and genomic medicine 10.64898/2026.08.26.26361466 medRxiv
Top 0.1%
4.3%
Show abstract

Unmeasured confounding threatens causal inference and replicability in observational multi-omic studies across variable environments. Genetic instrumental variables (Mendelian randomization) and negative-control calibration each address complementary sources of unmeasured confounding, yet no existing framework unifies them for omics-scale mediation analysis. We introduce ICONIC, an R package that embeds genetic instruments and negative controls within a proximal causal inference framework for total-effect and mediation analysis. ICONIC implements eight estimators spanning five confounding-control strategies, supports continuous, binary, and time-to-event outcomes, and provides extensive diagnostics including sensitivity analyses that map estimator performance across plausible assumptions. Ground-truth benchmarks are calibrated to real-omics covariance structures via a hybrid generative model (GAN + feature-level Gaussian copula) rather than parametric simulation, and a companion planning tool predicts performance gains from collecting additional omic data. We demonstrate ICONIC in two case studies: identifying placental transcriptomic mediators of gestational diabetes on birth weight (n = 164), and tumor-expression mediators of smoking intensity on lung cancer survival (n = 494). Notably, ICONIC's diagnostics recommended different estimation strategies across the two scenarios, reflecting differences in the likely influence of unmeasured confounding. ICONIC is freely available at https://github.com/sbresnahan/iconic/.

4
Genetic Architecture and Sample Size Impact Relative Performance of Nonlinear Machine Learning and Standard Polygenic Risk Scores

Zhu, J.; Baousi, A.; Morris, A. P.; Guo, H.

2026-09-03 genetic and genomic medicine 10.64898/2026.08.29.26361109 medRxiv
Top 0.2%
3.2%
Show abstract

Standard polygenic risk scores (PRSs) are constructed based on additive genome-wide association study (GWAS) summary statistics. Nonlinear machine learning methods have been increasingly applied to construct PRSs directly from individual-level data, with the aim of improving predictive performance over standard PRSs through their ability to model non-additive genetic effects. However, their superiority across studies has been inconsistent, and the conditions under which they provide meaningful improvements remain unclear. We combined theoretical analysis, simulations and a real-world application to investigate when two widely used nonlinear machine learning methods, random forest and XGBoost, outperform standard PRSs. Theoretical analysis showed that standard PRSs can implicitly capture part of the genetic variance attributable to nonadditive genetic effects through their contributions to marginal SNP effects, thereby losing less information than commonly assumed. Although nonlinear models have a higher theoretical potential, their greater flexibility incurs a bias-variance trade-off that can limit predictive gains at finite sample sizes. Simulations showed that XGBoost outperformed the standard PRS only when the genetic architecture involves a sufficiently large proportion of interaction genetic variance concentrated across relatively few interaction effects and large training samples were available. Random forest consistently underperformed the standard PRS. In an application to ischemic heart disease prediction using UK Biobank data, XGBoost showed no meaningful improvement in predictive performance over the standard PRS, whereas random forest again performed worse. Together, these findings suggest that nonlinear machine learning do not uniformly outperform standard PRSs; rather, their relative performance depends jointly on genetic architecture and training sample size. Our study helps to reconcile the inconsistent results reported across previous studies and provides a framework for identifying settings in which more complex PRS models are likely to be beneficial.

5
Primary Care Quality and Inappropriate Community Antibiotic Use: A Double Machine Learning Instrumental Variable Approach

Chen, Y.; Yi, H.; Rao, S.; Weber, A.; Hassmiller-Lich, K.; Sylvia, S.

2026-08-31 health economics 10.64898/2026.08.26.26361459 medRxiv
Top 0.3%
2.1%
Show abstract

Inappropriate antibiotic use presents a major global health challenge, particularly in low-resource settings where access to quality care is limited but antibiotics remain relatively unrestricted. This study estimates the causal effect of frontline primary care quality on inappropriate community antibiotic use, combining detailed community-based data from approximately 100 rural villages in rural China with an instrumental variable (IV) approach embedded within a double/debiased machine learning (DML) framework. We linked objective measures of village doctor clinical practice quality, measured through unannounced standardized patient visits, to household-level antibiotic use data collected from the same villages. To identify the causal effect, we constructed multiple candidate instruments from extensive provider characteristics and used an ensemble of machine learning algorithms within a flexible DML-IV framework to approximate an optimal instrument, addressing a many-weak-instruments problem. We found that improving village provider clinical practice quality reduced both antibiotic receipt during healthcare encounters for common diseases and household antibiotic storage for future self-medication. Our findings suggest that strengthening frontline primary care quality can meaningfully reduce inappropriate community antibiotic use without restricting access to essential treatment. More broadly, this study illustrates how causal machine learning can strengthen conventional causal estimation in complex observational settings in global health economics research.

6
Addressing Measurement Error of Machine-Learned Physical Activity in Nonlinear Dose-Response Survival Analysis: Development and Evaluation of Accelerated Failure Time, Spline, and Simulation-Extrapolation Method

Mamiya, H.; Zhang, Q.; Zhang, X.; Yan, Y.; Sharma, A.

2026-08-31 epidemiology 10.64898/2026.08.25.26361155 medRxiv
Top 0.4%
1.2%
Show abstract

Wearable (accelerometer) data and machine-learning allow objective assessment of the amount of daily physical activity. However, wearable-derived human activity is subject to measurement error. No studies have corrected the dose-response association between physical activity and survival time to chronic diseases, including cardiovascular disease (CVD). The objective is to estimate the measurement error-corrected association between CVD events and multiple measures of daily duration of light and total physical activity, derived from machine-learning and conventional accelerometer-processing methods. Our method combined an accelerated failure time model, spline, and simulation-extrapolation (SIMEX). The method recovered the true dose-response non-linear association in simulated data, while the naive model failed to capture it due to substantial attenuation. Application to the UK Biobank accelerometer cohort also showed an increased protective association of total physical activity after SIMEX correction (Time Ratio [TR] = 1.56, 95% CI: 1.28-1.82 vs. TR = 1.38, 95% CI: 1.24-1.54 for SIMEX-corrected vs. uncorrected dose-response association between the 95th and 5th percentiles of total activity), with a similar increase for light physical activity. Sensitivity analysis indicates that the female population experiences a substantially larger protective association after SIMEX correction than males. Dose-response survival analysis is a widely used analytical method in physical activity epidemiology and benefits from measurement error correction.

7
New tests for trials of very few patients using longitudinal data - a case-study in Autosomal Recessive Cerebellar Ataxias

Hendrickx, N.; Mentre, F.; Karlsson, M. O.; Hooker, A. C.; Traschütz, A.; Schüle, R.; PROSPAX Consortium, ; EVIDENCE-RND Consortium, ; Synofzik, M.; Comets, E.

2026-09-02 health informatics 10.64898/2026.08.28.26361588 medRxiv
Top 0.5%
0.9%
Show abstract

We propose two new tests to detect drug effects (DE) in trials of one to very few patients followed during two periods (before and after initiation of a treatment). Both methods use longitudinal natural history data to inform the estimation of each patient's DE. The first method uses a non linear mixed effect model (NLMEM) reflecting an expected natural history with a hypothetical drug effect, to estimate the Conditional Distribution of the Drug Effect (CDDE). The second method trains a Pareto Depth Analysis (PDA) algorithm, a machine learning based approach based on outlier detection, that we implement using data simulated under the NLMEM. We evaluated the two tests with a simulation study. We used data from the PROSPAX study in Autosomal Recessive Cerebellar Ataxias (ARCAs, to derive a NLMEM for the Scale for the Assessment and Rating of Ataxia score. The CDDE method provided controlled type I error and, in some scenarios, adequate corrected power, though sensitivity analyses showed vulnerability to misspecification. The PDA method demonstrated lower statistical power except with high score precision. These results highlight different strategies for quantifying treatment effects in ultra rare, patient' specific trials. They can inform methodological design for future ARCA precision therapies.

8
Projected Population-Level Impact of Digital Return of Results for Cardiovascular-Kidney-Metabolic Screening at US Blood Donation Centers: A Monte Carlo Simulation Study

Qian, Z.; Khera, A.; Makhnoon, S.; Chapman, B. E.; Bryant, B.; Sayers, M.; Compton, F.; Eason, S.; Xing, C.; Ahmad, Z.

2026-09-03 public and global health 10.64898/2026.09.01.26360806 medRxiv
Top 0.5%
0.9%
Show abstract

Background. Cardiovascular-kidney-metabolic (CKM) syndrome affects nearly 90% of US adults, yet most individuals at early, modifiable stages remain unidentified outside clinical care. Blood donation centers offer a scalable, non-clinical venue for CKM screening, but the potential benefit of screening in this context remains unclear. We projected the population-level impact of effective digital return of results (ROR) to inform the design of a pragmatic trial. Methods. We developed a Monte Carlo simulation (100,000 iterations) of the incident major adverse cardiovascular events (MACE), end-stage renal disease (ESRD), and type 2 diabetes (T2DM) preventable by ROR-prompted, guideline-concordant follow-up among donors in CKM Stages 1-2. The estimand counts only events averted by donors who act because of ROR; the intervention effect was modeled directly on strictly positive support, and action was translated into prevented events through a hazard-based cumulative-incidence difference that counts each donor at most once. We evaluated 18 design cells (donor volumes 300,000, 1 million, and 8 million/year; 5- and 10-year horizons; action-rate gains of +10, +20, and +30 percentage points [pp]) and, in a complementary two-arm simulation, the assurance (expected power) of detecting the effect in a single deployment. Results. Under the primary +20 pp scenario, ROR at a single large blood center (300,000 donors/year) is projected to prevent a median of 2,201 events (95% uncertainty interval [UI], 1,099-4,364) over 10 years, scaling to 58,526 (29,154-116,769) at the national donor pool. All 18 design cells had strictly positive 95% lower bounds. The number needed to screen was 136 and the screening cost $2,045 per event prevented (at $15/donor), both invariant to donor volume. Impact scaled linearly with volume and effect size but sub-linearly with the horizon. Detection of the effect was effectively certain at gains of +20 pp or larger (assurance [≥]99.6% in every cell and >99.9% in all but the smallest 5-year cell). Conclusions. Even under the conservative scenario, digital CKM ROR at blood donation centers is projected to prevent hundreds to tens of thousands of incident cardiometabolic events at a screening cost per event well within accepted prevention benchmarks, providing prospective, quantitative justification for a pragmatic, randomized evaluation of digital ROR in non-clinical screening settings.

9
Can Dental AI Really Beat Dentists? DentalPair-Cert for Rigorous AI-Dentist Inference

Alve, S. R.; Rahman, S.; Meem, S. M. A. C.

2026-09-02 dentistry and oral medicine 10.64898/2026.09.01.26361874 medRxiv
Top 0.8%
0.5%
Show abstract

A dental AI system and a dentist reading the same radiographs form a paired comparison. Published comparative studies often report the two arms separately against a reference standard, leaving the joint pattern of correctness between them unavailable for secondary paired inference. We show what that omission costs. The accuracy difference remains exactly identified; its sampling variance does not, so the report contains the estimate and not its uncertainty. On a study of 282 units, two published accuracies are consistent with 38 distinct joint tables whose confidence intervals differ in width by a factor of 2.5. The consequence is a three-zone decision map rather than a single threshold: differences at or below 1.06 points are non-significant under every compatible table, differences at or above 6.03 points are significant under every compatible table, and in between the published numbers cannot decide. We then show the omission is repairable at negligible cost. One additional integer, the number of units both arms classify correctly, identifies the joint table exactly and restores standard paired inference. For a panel of readers the pairwise dependences must arise from one joint distribution, a constraint that binds once three readers are present; publishing each reader's joint-correct count against a single reference reader cannot widen and may tighten every pairwise bound, and in a 7-arm experiment reduced them by a median of 37% even for pairs excluding that reference. Where the integer was never published we give DentalPair-Cert, an interval with finite-sample coverage uniformly over every admissible within-unit AI-dentist dependence under the independent-sampling-unit model, certified in both the nuisance maximization and the inversion. Across 4,200,000 simulated comparisons an independence analysis falls to 74.5% coverage with 12.2% type-I error; in a purposive sample of 9 recent comparative studies, 1 reported a paired test on discordant units.

10
Multi-season evaluation and analysis of categorical trend forecasts of influenza hospital admissions in the United States

Davis, J. T.; Kaur, G.; Hines, A.; Ben-Nun, M.; Venkatramanan, S.; Brooks, L.; Mathis, S.; Ajelli, M.; Litvinova, M.; Kummer, A. G.; Ventura, P. C.; Mhade, S.; Weber, D.; Shemetov, D.; DeFries, N.; McDonald, D. J.; Yamana, T.; Zepeda-Tello, R.; Shaman, J.; Yaari, R.; Pei, S.; Webber, A.; Shandross, L.; Ray, E.; Wadsworth, S.; Niemi, J.; Redman, W. T.; Mullany, L.; Posner, R.; Mallela, A.; Lin, Y. T.; Hlavacek, W. S.; Smart, A.; Gill, A. A.; Drennan, A.; Fiebiger, B. J.; Miller, E. F.; Lee, J.; Mihaljevic, J. R.; Geist, K. A.; Baltz, M.; Bernik, O.; Truong, Y.-M. B.; Chen, Y.; Grosvenor, C. J.;

2026-09-02 epidemiology 10.64898/2026.08.31.26361843 medRxiv
Top 0.9%
0.4%
Show abstract

Forecasting influenza hospitalizations informs public health preparedness, yet questions remain about which types of forecasts best guide action. We evaluate categorical trend forecasts, which communicate probabilities of upcoming increases or decreases in epidemic trajectories, submitted to CDC's FluSight Forecasting Challenge between Fall-2024 and Spring-2026. Teams submitted probability distributions over five categories describing direction and magnitude of week-over-week changes in laboratory-confirmed influenza hospital admissions. We assessed performance using Ranked Probability Skill Score, Brier Skill Score, and measures of forecast-observation agreement. Most models outperformed an equal-probability baseline; the FluSight ensemble ranked among the top three in the 2024-25 and 2025-26 seasons. Forecasts were most accurate during stable periods and least during periods of rapid change, with most models underestimating observed trends. Conclusions were robust to choice of scoring metric and reference model. These results support categorical trend ensembles as an approach to communicating infectious disease forecasts that may inform public health decision-making.

11
Characterizing the Burden of Scrub Typhus in Nepalese Children using Representative School-Based Cross-Sectional Serosurveys

Naga, S. R.; Shahi, S. B.; Gosai, S.; Maharjan, M.; Katuwal, N.; Morrison, E.; Andrews, J.; Shrestha, R.; Tamrakar, D.; Aiemjoy, K.

2026-08-31 infectious diseases 10.64898/2026.08.28.26361601 medRxiv
Top 1%
0.3%
Show abstract

Background Scrub typhus, an acute bacterial infection caused by Orientia tsutsugamushi, is an important, under-recognized etiology of febrile illness in South Asia. Accurate surveillance for scrub typhus in Nepal is challenging due to non-specific symptoms and limited diagnostic tools. Methods We conducted a representative school-based cross-sectional serosurvey in Kavrepalanchok and Dolakha, Nepal between November 2021 and April 2022 using two-stage sampling, we randomly selected 13 public schools and then up to 100 children aged 4-18 years per school. We collected capillary blood samples and tested for IgG responses to Orientia tsutsugamushi-derived recombinant 56-kDa antigen using commercially available ELISA kits. We estimated seroincidence rates using previously-published models of antibody decay dynamics from confirmed scrub typhus cases. We compared seroincidence to seroprevalence estimates using cutoffs derived from Gaussian finite mixture models applied to the study population. Results We enrolled a total of 827 children (participation rate: 94.8%). The median age was 10 years (IQR: 8-13), and 53.08% of the participants were female. The overall seroincidence rate was 6.3 new infections per 1000 person-years (95% CI: 4.7- 8.4), and the rate was slightly higher in peri-urban Kavrepalanchok (7.3; 95% CI: 4.9 - 10.3) compared with rural Dolakha (5.3; 95% CI: 3.3 - 8.5). Seroincidence increased with age, from 3.6 new infections per 1000 person-years (95% CI: 1.4- 9.7) among children aged 4- 7 years to 9.4 (95% CI: 6.2-14.3) among children aged 14-18 years. Seroincidence was higher in females (8.3; 95% CI: 5.9 -11.8) than in males (4.1; 95% CI: 2.4 - 6.9) (p=0.02). The overall seroprevalence was 6.5% (95% CI: 4.9-8.4) and followed a similar age and geographic trend to the seroincidence rate. Discussion Our findings reveal a substantial burden of pediatric scrub typhus in the Kavrepalanchok and Dolakha districts of Nepal. Seroincidence increased with age and was higher among females. School-based serosurveys offer an efficient sampling frame to rapidly assess population-level scrub typhus transmission intensity among children and adolescents, though may not represent out-of-school populations.

12
Heterogeneity in pre-vaccination population immunity can contribute to variability in vaccine effectiveness estimates

Pillai, A. N.; Park, S. W.; Lipsitch, M.; Cowling, B. J.; Cobey, S.

2026-08-31 epidemiology 10.64898/2026.08.29.26361716 medRxiv
Top 2%
0.1%
Show abstract

Vaccine effectiveness (VE) estimates can vary widely between years and populations, even for the same vaccine. Estimated VE is known to be sensitive to susceptible depletion and differences in pre-vaccination infection risk between vaccinated and unvaccinated populations. However, how variation in pre-vaccination risk within and between the two groups affects VE estimates over time remains unclear. This uncertainty is especially important given negative VE estimates. We investigated the difference between estimated VE and true vaccine protection considering continuous distributions of pre-vaccination infection risk under three scenarios. When the vaccinated and unvaccinated populations differ in their mean risk, estimated VE can be higher or lower than true vaccine protection. Similar patterns arise when both populations share identical means but different risk distributions. Finally, if infection-derived immunity lasts longer than vaccine protection, annual VE estimates can vary by tens of percentage points between years despite constant true vaccine protection. These theoretical results underscore that VE studies estimate contrasting risk between vaccinated and unvaccinated individuals in a particular time and place, and VE estimates can vary counterintuitively between years and populations even with constant vaccine-induced protection. Explaining variability in estimated VE thus requires a more complete understanding of populations' distributions of infection risk.

13
RedFuMOS: A novel approach for multi-omics and clinical data-driven patient stratification

De Luca, S.; Fava, C.; Rizzo, G.; Visconti, A.; Berchialla, P.

2026-08-31 health informatics 10.64898/2026.08.26.26361415 medRxiv
Top 2%
0.1%
Show abstract

Background. Patient stratification from multi-omics and clinical data is essential for uncovering disease heterogeneity and moving toward more personalized treatment strategies. However, integrating heterogeneous data layers while identifying robust patient strata remains challenging. Methods. We introduce Reduced Fusion of Multi-Omics Stratification (RedFuMOS), a novel three-step approach for patient stratification based on mixed-type multi-omics data. RedFuMOS extends Similarity Network Fusion to accommodate mixed-type data layers and layer-specific similarity measures for data integration, includes a dimensionality reduction step to mitigate the curse of dimensionality, and performs patient stratification using density-based hierarchical clustering with HDBSCAN. It also implemented an automated optimization procedure to identify the best set of hyperparameters, minimizing the need for manual tuning. Results. RedFuMOS outperformed six state-of-the-art tools for multi-omics patient stratification in a comprehensive simulated benchmarking study, which also confirmed that, although computationally expensive, the dimensionality reduction step is crucial for achieving good stratification performance. Additionally, RedFuMOS identified two clinically relevant patient strata in a small real-world cohort of patients with Philadelphia chromosome-positive chronic myeloid leukaemia. Conclusion. RedFuMOS provides a flexible framework for integrating heterogeneous multi-omics and clinical data. RedFuMOS is available as an R package at http://github.com/delucasara/RedFuMOS.

14
Machine Learning-Based Prediction of Maternal Morbidity across Heterogeneous Populations in the United States using Sequential Modeling of the All of Us Dataset

Zhuang, H.; Zakama, A.; Heller, K.; Faulkner, S.; Gollub, B.; Young-Lin, N.; Chen, I. Y.; Asiedu, M.

2026-08-31 obstetrics and gynecology 10.64898/2026.08.25.26360552 medRxiv
Top 3%
0.1%
Show abstract

In this work, we demonstrate the unprecedented value of NIH's "All of Us Research Program" (AoURP) dataset in studying maternal morbidity and building predictive machine learning (ML) models across heterogeneous populations in the United States. We developed robust and data-driven preprocessing pipelines to curate a longitudinal, multi-site, multimodal, and demographically diverse pregnancy dataset (20,253 subjects; 27,525 pregnancy episodes) from AoURP data, using electronic health records (EHR) (Conditions, Labs, Measurements) and survey responses (Social Determinant of Health (SDoH)), focusing on 7 crucial maternal health adverse outcomes. After characterizing data quality, missingness, and heterogeneity, we performed statistical correlation analysis to identify risk factors. We subsequently developed XGBoost and sequential LSTM models to predict the adverse outcomes, reaching state-of-the-art performance for multiple outcomes. We conducted model interpretability post-hoc analysis to understand success points and fairness analysis to evaluate implications for socio-economic disparities. Four practicing physicians reviewed the set of statistically significant and ML model identified features to assess their clinical validity and novelty. Most features identified through either statistical correlations or ML feature importance analysis aligned with known clinical risk factors. Several features were identified that the ML models used but that are not currently used in clinical practice and may merit further clinical investigation. Fairness analysis revealed certain associations with SDoH and age highlight areas that warrant continued monitoring. Overall, we demonstrate that meaningful populational level patterns can be extracted, and high-performing machine learning models can be trained on this longitudinal, diverse, multi-site dataset. Important risk features, particularly novel ones identified, if validated, could inform new strategies for maternal care or enable development and validation of outcome-specific, clinically deployable ML models.

15
A mechanistic statistical model of dengue dynamics in an endemic region

Luna-Martinez, N.; Cruz-Rodriguez, E. X.; Bernal-Castro, E. A.

2026-09-03 epidemiology 10.64898/2026.09.01.26361961 medRxiv
Top 3%
0.1%
Show abstract

Background Dengue is a major public health challenge, and predictive models are crucial for early warning systems. However, many current modeling practices rely exclusively on climatic factors or employ complex algorithms that lack the interpretability needed for informed public health decision-making. To address these shortcomings, we developed and validated a multidimensional, interpretable statistical model to predict monthly dengue incidence. Methodology/Principal Findings We used a Generalized Linear Mixed Model (GLMM) with a Negative Binomial distribution to analyze 14 years (2010-2023) of spatiotemporal data from 37 municipalities in Huila, Colombia, an endemic region. The model integrates non-linear and lagged effects of climatic, demographic, and socioeconomic factors. The final model underwent rigorous external validation on an independent test set (2021-2023). Our model demonstrated high predictive discrimination (R2 = 0.743, Spearman's {rho} = 0.657), accurately capturing the timing of epidemic outbreaks. Key findings include the identification of an optimal thermal window for transmission at 27-28{degrees}C, a threshold effect for precipitation above 800 mm, and a saturation dynamic in outbreak autocorrelation. Conclusions/Significance This mechanistically-informed statistical approach provides a robust and transparent tool for epidemiological surveillance, successfully balancing high predictive performance with the explanatory power needed for effective, data-driven public health interventions.

16
Clinical evaluation of artificial intelligence for diagnostics of antibiotic-resistant bacteria

Hessel, M.; Inda Diaz, J. S.; Sjöberg, A.; Salva-Serra, F.; Helldal, L.; Jirstrand, M.; Johnning, A.; Kristiansson, E.; Skovbjerg, S.

2026-08-31 infectious diseases 10.64898/2026.08.27.26361401 medRxiv
Top 3%
0.1%
Show abstract

Antimicrobial resistance is a public health challenge, driving the need for rapid, cost-effective diagnostic support tools. Artificial intelligence (AI) may enable prediction of susceptibility to untested antibiotics from known susceptibility results, but prospective clinical validation is required before routine use. We evaluated an AI-based decision support method, trained on invasive isolates from the European Surveillance System (TESSy), for prediction of antibiotic susceptibility in clinical Escherichia coli urine isolates. The evaluation included 99 E. coli isolates from urine samples with diversity in age, sex, and antibiotic susceptibility. Predictions were evaluated for 14 antibiotics using patient metadata and susceptibility results for 4-8 antibiotics as input. Prediction uncertainty was handled using conformal prediction, allowing abstention when confidence was insufficient. EUCAST disk diffusion test results were used as reference and genomic sequence data was used to explore mechanisms of the AI performance. Without conformal prediction, 84% of predictions were correct when susceptibility results of six antibiotics were used to predict susceptibility to eight additional antibiotics. Across all predictions generated using susceptibility results for six antibiotics as input, the major and very major error rates were 19% and 12%, respectively. Prediction errors varied between antibiotics and were associated with certain phenotypic and genotypic resistance patterns. Conformal prediction reduced errors but increased abstentions; at confidence levels of 90%, 95%, and 97.5%, the model abstained in 9.6%, 14%, and 22% of instances. The method showed promising performance, but its clinical use remains limited and may require diagnostic data beyond susceptibility test results and demographic variables.

17
Understanding RSV Resurgence Following COVID-19 in Ontario, Canada: Evaluating the Roles of Contact Patterns and Maternal Immunity

Parpia, A.; Wright, J.; Gharouni, A.; Thampi, N.; Fitzpatrick, T.

2026-08-31 epidemiology 10.64898/2026.08.28.26361657 medRxiv
Top 3%
0.1%
Show abstract

Background: Respiratory syncytial virus (RSV) remains a leading cause of hospitalization in infancy, with severe outcomes influenced by both contact patterns and passive immunity. Non-pharmaceutical interventions (NPIs) during the COVID-19 pandemic suppressed RSV circulation and reduced opportunities for maternal immune boosting, potentially altering protection among newborns. We evaluated whether incorporating time-varying maternal immunity improves the ability of an age-structured transmission model to predict post-pandemic RSV hospitalization patterns in infants. Methods: We analyzed population-based RSV hospitalizations among Ontario (Canada) infants (<1 year) from July 2, 2017 to June 25, 2024, using linked administrative databases. A deterministic compartmental model across seven age classes was calibrated against pre-pandemic data using Latin Hypercube Sampling. We compared a model incorporating time-varying contact rates alone against a specification that additionally included time-varying maternal immunity. Results: Both specifications accurately reproduced pre-pandemic seasonality and macro-level post-pandemic resurgence features. The constant maternal immunity model showed slightly better accuracy in capturing the 2021/22 peak compared to the time-varying maternal immunity specification. However, both qualitatively captured the continued near-absence of RSV and the observed peak was captured within the 95% credible intervals. While both models precisely captured the timing and overwhelming surge of admissions that occurred in 2022/23, they failed to capture the premature peak timing and magnitude in 2023/24. Conclusions: Incorporating time-varying maternal immunity did not improve model accuracy post-pandemic. While maternal protection is essential for evaluating infant immunizations, population-level contact shifts primarily shaped post-pandemic RSV seasonality, indicating that models must account for these mechanisms of RSV transmission dynamics.

18
Prospective In-silico Simulation of the VESALIUS-CV Trial Using Biomedical Knowledge Graph and Real-World Data-Driven AI Modeling

Perlman, A.; Goldstein, N.; Goldman, M.; Shapiro, M.; Barash, E.; Bar, A.; Raveh, T.; Tordjman, E.; Schussheim, H.; Dormont, F.; Matalon, O.

2026-08-31 cardiovascular medicine 10.64898/2026.08.26.26361436 medRxiv
Top 3%
0.1%
Show abstract

Background. Cardiovascular-outcomes trials are lengthy, costly, and associated with substantial uncertainty prior to readout. In-silico trial simulation using real-world data (RWD) has emerged as a potential tool to support earlier decision-making; however, evidence of prospective predictive validity, generated prior to trial result disclosure, remains limited. Methods. We applied a semi-mechanistic machine learning framework integrating real-world patient data with biologically informed drug representations to prospectively simulate the VESALIUS-CV trial evaluating evolocumab versus placebo. The simulation model was trained on a combination of patient-level real-world data and a drug-centric knowledge graph and validated for both patient-level and trial-level retrospective predictive performance. The model was then used to simulate VESALIUS-CV before public disclosure of trial results, using a locked model and prespecified eligibility criteria and primary endpoint aligned with the clinical protocol. A patient-level time-to-event model was used to generate virtual trial arms, from which cumulative incidence curves, hazard ratios, confidence intervals, and p-values for major adverse cardiovascular events (MACE) were estimated. Results. In retrospective validation, the model demonstrated strong patient-level discrimination, with time-dependent ROC-AUC values ranging from 0.80 to 0.90 across follow-up horizons. For trial-level validation, 22 randomized cardiovascular-outcomes trials were simulated, and hazard ratios for 3-point MACE across 24 between-arm comparisons showed consistent directional agreement and quantitative correlation with published results such that the model accurately predicted trial success, achieving an F1 score of 0.83, with precision of 0.79 and sensitivity of 0.89. In a fully prospective application, the simulation predicted a statistically significant reduction in 3-point MACE with evolocumab versus placebo, estimating a hazard ratio of 0.78 (95% CI, 0.70-0.87) at 54 months. These predictions were consistent with the subsequently reported VESALIUS-CV results, which demonstrated a hazard ratio of 0.75 (95% CI, 0.65-0.86) at 55 months of median follow-up. Conclusions. In a fully prospective setting, a RWD-driven, AI-based simulation accurately predicted the direction, magnitude, and temporal dynamics of treatment effects observed in the VESALIUS-CV trial. These results demonstrate that in-silico trial simulation can anticipate clinical outcomes in the prospective setting, supporting its use as a complementary tool for early decision-making, trial design optimization, and de-risking in cardiovascular drug development.

19
Associations of the Patient Safety Screener-3 With Depression and Suicide Risk: A Nationwide Cross-Sectional Study in Japan

Kiryu, K.; Tamune, H.; Takahashi, K.; Fujikawa, H.; Harada, H.; Fukui, S.; Nagasaki, K.; Nishizaki, Y.; Kato, T.; Tokuda, Y.

2026-08-31 psychiatry and clinical psychology 10.64898/2026.08.30.26361711 medRxiv
Top 3%
0.0%
Show abstract

Aim: The Patient Safety Screener-3 (PSS-3) is a brief suicide-risk screening tool. Item 1 of this scale assesses depressive mood but is not included in the total score. We examined the association of item 1 with depressive symptom severity and characterized the suicide-related risk captured by PSS-3 total positivity. Methods: We conducted a nationwide cross-sectional survey among resident physicians in Japan. Associations between PSS-3 item 1 endorsement and Patient Health Questionnaire-9 (PHQ-9) scores were evaluated using the Wilcoxon rank-sum test. Diagnostic performance of item 1 was evaluated using PHQ-9 positivity ([&ge;]10) as reference standard. We also compared Short-form Scale for Suicide Ideation (SIS-6) scores according to PSS-3 total positivity and PHQ-9 item 9 positivity. Results: A total of 1,844 participants were included. PSS-3 item 1 was endorsed by 443 physicians (24.0%), and 47 (2.5%) met the criteria for PSS-3 total positivity. Item 1 showed 79.3% sensitivity and 79.5% specificity for PHQ-9 positivity. SIS-6 scores were higher in the PSS-3 total-positive group than in the total-negative group (median [IQR], 6 [5-9] vs 0 [0-1]; p<0.001). The SIS-6 showed a higher area under the receiver operating characteristic curve (AUC) and Youden index using PSS-3 total positivity (AUC, 0.961; optimal cutoff, 3) than PHQ-9 item 9 positivity (AUC, 0.907; optimal cutoff, 2). Discussion: PSS-3 may support brief, simultaneous screening for depressive symptoms and suicide-related risk. Compared with PHQ-9 item 9, PSS-3 may capture a more severe spectrum of suicide-related risk. PSS-3 may facilitate identification of individuals requiring further mental health assessment.

20
Burr-Hole Intersection of Middle Meningeal Artery Branches and Recurrence in Chronic Subdural Haematoma: a Multicentre Retrospective Cohort Study

Saba, T. M.; Moudgil-Joshi, J.; Pandit, A. S.; Penn, J.; Mallon, D.; Marcus, H. J.; Grover, P.

2026-08-31 surgery 10.64898/2026.08.26.26361348 medRxiv
Top 3%
0.0%
Show abstract

Background and Objectives: Recurrence following burr-hole drainage of chronic subdural haematoma (cSDH) occurs in 10-25% of cases, sustained by neovascularisation of the subdural neomembrane supplied by the middle meningeal artery (MMA). MMA embolisation reduces recurrence; whether incidental burr-hole intersection of MMA branches during drainage confers similar benefit is unknown. Methods: We performed a multicentre retrospective cohort study of consecutive adults undergoing burr-hole drainage for cSDH at two UK tertiary neurosurgical centres. Postoperative thin-slice CT was used to classify burr-hole intersection of the underlying MMA groove (no hit, distal-branch hit or main-branch hit) and measure perpendicular burr-hole-to-MMA-groove distance. Co-primary outcomes were radiological recurrence and recurrence requiring intervention. Patient-clustered multivariable logistic regression adjusted for prespecified clinical covariates and treating site. Results: 227 patients (284 operated hemispheres) were included. Radiological recurrence decreased from 34.4% with no branch hit to 22.9% with main-branch intersection, with the gradient confined predominantly to unilateral cSDH. Main-branch intersection was associated with lower adjusted odds of radiological recurrence in unilateral cSDH (adjusted OR 0.30, 95% CI 0.11- 0.81; P = .018), with a similar but non-significant association in the overall cohort (adjusted OR 0.53, 95% CI 0.26-1.07; P = .075). Burr-hole-to-MMA-groove distance demonstrated a more consistent association: in the overall cohort, each 5-mm increase independently increased the odds of radiological recurrence (adjusted OR 1.38, 95% CI 1.04-1.82; P = .025). In unilateral cSDH, each 5-mm increase was independently associated with both radiological recurrence (adjusted OR 1.45, 95% CI 1.03-2.04; P = .034) and recurrence requiring intervention (adjusted OR 1.52, 95% CI 1.05-2.20; P = .027). Conclusion: Main-branch intersection of the middle meningeal artery during routine burr-hole surgery is associated with lower recurrence of unilateral cSDH, while the accompanying burr-hole-to-MMA-groove distance gradient provides biologically plausible support for a dose-response relationship. Together, these findings provide mechanistic rationale for prospective evaluation of intentional neuronavigation-guided MMA targeting (BURR-MMA; NCT07549893).